Back

The Plant Genome

Wiley

Preprints posted in the last 90 days, ranked by how well they match The Plant Genome's content profile, based on 57 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.

1
Phenotyping replication is a major determinant of genomic predictive ability in sweet sorghum (Sorghum bicolor Moench)

CHARLES, J. R.; Rice, B.; Tovignan, T.; Morris, G. P.; Pressoir, G.

2026-06-19 genomics 10.64898/2026.06.15.731123 medRxiv
Top 0.1%
35.4%
Show abstract

Genomic selection can increase the rate of genetic gain in crop breeding programs, but its effectiveness depends on the reliability of phenotypic data, the size and composition of the training population (TP), and the statistical model used to estimate genomic breeding values. These design choices are especially important in resource-limited breeding programs, where additional replication, larger TPs, and more extensive genotyping compete for the same resources. Using empirical data from a sweet sorghum [Sorghum bicolor (L.) Moench] breeding population, developed by CHIBAS, we evaluated the effects of phenotyping replication, TP size, training-validation genomic relatedness, and genomic prediction (GP) model on predictive ability (PA). Grain yield, plant height, stem weight, and total soluble solids were evaluated across three field environments. Few studies in sorghum have examined these factors together with comparable empirical rigor. Increasing replication improved genomic heritability and PA for all traits and environments, with the largest gains observed for grain yield. Larger TPs and increased training-validation genomic relatedness also improved PA, but their effects were most significant when phenotype estimates were based on multiple replicates. GP models showed largely comparable PAs across all evaluated traits. Different models produced similar PA, with a few exceptions. These findings provide practical guidance for optimizing genomic selection in resource-limited sorghum breeding programs. ARTICLE SUMMARYGenomic selection can accelerate breeding only when the phenotypes used to train prediction models have high reliability. Using a sweet sorghum breeding population evaluated in three Haitian field environments, we quantified how replication number, training population size, training-validation genomic relatedness, and prediction model affected genomic predictive ability for grain yield, plant height, stem weight, and total soluble solids. Replication increased genomic heritability and predictive ability for all traits, with the strongest effects for grain yield. Larger and more connected training populations improved prediction, mainly when replication was adequate. These results provide practical guidance for resource-limited breeding programs. Core ideasO_LIIn this empirical sweet sorghum breeding population, phenotyping replication was the dominant factor explaining variation in genomic predictive ability across traits and environments. C_LIO_LIThe benefit of larger training populations and greater training-validation genomic relatedness increased when phenotype estimates were based on more replicates. C_LIO_LIGrain yield, the most environmentally sensitive trait evaluated, showed the largest response to improved replication and training-population design. C_LIO_LIBayesian models, rrBLUP, and GBLUP showed similar predictive abilities across traits and environments, suggesting that phenotyping and experimental design may be more important than model complexity. C_LI

2
Integrating Genome-Wide Association Analysis, Functional Annotation and Regulatory Genomics for the Prioritization of Candidate Variants Associated with Grain Yield in Upland Rice

da Cruz, A. C.; Vianello, R. P.; Valdisser, P. A. M. R.; Bueno, L. G.; Brondani, C.

2026-07-16 genomics 10.64898/2026.07.10.737825 medRxiv
Top 0.1%
28.0%
Show abstract

Grain yield is a highly complex quantitative trait in rice, resulting from the interaction of multiple genetic, physiological and environmental factors. Although genome-wide association studies (GWAS) have successfully identified loci associated with grain yield, translating statistical associations into biologically meaningful candidate variants remains a major challenge, particularly for variants located in regulatory regions. This study aimed to identify genomic variants associated with grain yield in upland rice and to develop an integrative framework for functionally prioritizing candidate variants through the combination of genome-wide association analysis, functional annotation and regulatory genomics. A panel of 252 accessions from the Brazilian Rice Core Collection was phenotyped for grain yield and genotyped with 35,763 single nucleotide polymorphism (SNP) markers. GWAS identified 29 SNPs significantly associated with grain yield, including 16 variants located within or near annotated genes and 13 located in intergenic regions. The identified candidate genes were involved in signal perception, metabolite transport, amino acid and energy metabolism, hormone biosynthesis, protein turnover, RNA processing and disease resistance, highlighting the polygenic architecture of grain yield. Functional characterization of the intergenic regions revealed enrichment of cis-regulatory elements recognized by transcription factors associated with hormonal signaling, drought response, carbon metabolism, photosynthesis and reproductive development, indicating that regulatory variation represents an important component of grain yield determination. By integrating GWAS signals, candidate gene annotation, cis-regulatory element characterization and the physical proximity between SNPs and cis-regulatory elements, an integrative prioritization strategy identified seven intergenic SNPs as the most promising candidates for functional validation. Together, these findings establish a robust framework for discovering, prioritizing and functionally validating regulatory variants, bridging the gap between statistical associations and biological function while providing a rational strategy for translating GWAS discoveries into molecular breeding of complex traits.

3
CannSelect: A High-Quality Genotyping Platform for Cannabis sativa

Wilkerson, D. G.; Stack, G. M.; Carlson, C. H.; Quade, M. A.; Dowling, C. A.; Toth, J. A.; Murdock, M. J.; Jasinski, J.; Stansell, Z. J.; McKay, J. K.; Smart, L. B.

2026-08-21 genomics 10.64898/2026.08.18.745408 medRxiv
Top 0.1%
26.0%
Show abstract

The field of genomics has enabled extraordinary progress in horticultural crop research. However, there is still a need for cost-effective, high-resolution technologies flexible to the diversity found in emerging crops. To this end, we introduce CannSelect, a high-quality genotyping platform for Cannabis sativa. Designed for use in diversity analyses and trait mapping, probe targets were selected from four genotyped diversity panels and a curated gene list. This platform has been used to effectively map day-neutrality in a segregating population to the Autoflower1 locus with average capture efficiencies of 88.5%. With broad genome coverage, demonstrated target specificity, and reproducibility, CannSelect is expected to perform well across the diversity of C. sativa. We describe the methodology used to design CannSelect v1.0 and performance metrics for testing capture efficiency and target alignment in diverse genome assemblies. The CannSelect platform represents a robust and scalable, genome-wide genotyping tool for C. sativa researchers and breeders.

4
Enhancing predictive accuracy of yield traits in cassava through multi-trait genomic prediction

de Freitas, G. M.; Certuche, D. S.; Jannink, J.-L.; de Oliveira, E. J.; Garcia, A. A. F.

2026-07-06 genetics 10.64898/2026.07.01.735838 medRxiv
Top 0.1%
22.2%
Show abstract

Multi-trait genomic prediction offers a practical route to improve selection for costly, complex traits in clonally propagated crops such as cassava. In a Brazilian breeding panel of 1,078 cassava clones genotyped with 25,923 SNPs and phenotyped for six agronomic traits, we compared single-trait (ST) and multi-trait (MT) GBLUP models. Stage-wise mixed models produced BLUEs that fed into ST and MT-GBLUP. We tested five cross-validation schemes that mimic breeder realities: ST baseline (CV1); naive all-traits MT prediction for unphenotyped candidates (CV2); MT prediction using auxiliary trait phenotypes in the test set (CV3); and two sparse-phenotyping regimes with missingness by trait (CV4) or by clone (CV5) at 25%, 50%, and 75% levels. The main results were that, under the ST baseline (CV1), predictive ability ranged from 0.50 for DMC and 0.45 for FRY down to 0.13 for Le.Dis. A naive full MT model (CV2) performed approximately on par with ST-GBLUP. In contrast, MT designs (CV3) that included informative auxiliary traits, such as shoot yield and combinations with plant vigor and leaf disease severity, yielded small gains for DMC with predictive ability of approximately 0.51 (+2%), while FRY predictive ability increased to approximately 0.65 (+44%), accompanied by RMSE reductions for FRY up to approximately 13.5% (e.g. RMSE approximately 6.2). Sparse-phenotyping simulations (CV4/CV5) demonstrated that MT models sustain or even improve predictive ability under realistic missing-data regimes (PA {approx} 0.62 - 0.65). Selection concordance between MT and ST top-10% sets was generally high (>0.80), and MT configurations produced measurable improvements in expected selection response and genetic gain per cycle for several target traits. These results indicate that strategically implemented MT-GBLUP, using a small set of biologically and operationally informative auxiliary traits and optimized sparse phenotyping, can materially increase predictive accuracy and selection efciency for economically critical cassava traits while reducing phenotyping burden.

5
Near-infrared phenomic and genomic prediction for seed protein in winter legume white lupin (Lupinus albus L.): A utility comparison

Castillo, M. P.; Oyebode, O. G.; Lenahan, A.; Orloski, A.; Wolfe, M.

2026-08-11 genomics 10.64898/2026.08.05.743001 medRxiv
Top 0.1%
22.0%
Show abstract

White lupin (Lupinus albus L.) is a cool-season grain legume with seed crude protein of 33-47%, competitive with soybean (Glycine max L.) meal. It also fixes nitrogen and mobilizes soil phosphorus. Because soybean is a summer crop, white lupin can occupy Southeastern winter fields as a complementary protein source. Breeding for seed protein is limited by the cost and throughput of reference phenotyping. To determine how each is best deployed, we compared the utility of near-infrared spectroscopy (NIRS)-based phenomic selection with genomic selection based on 246,847 SNPs from low-pass, whole genome sequencing in a panel of Auburn University breeding lines and USDA National Plant Germplasm System germplasm. A handheld NIR calibration against Dumas reference protein reached screening-grade accuracy (R2 = 0.81). Under common cross-validation, phenomic predictive ability was 0.93 and genomic was 0.12. The low genomic value was consistent with moderate heritability (H2 = 0.33) and strong genotype-by-year interaction. Beyond predictive ability, NIRS recovered superior accessions the strictest selection intensity, and 40 to 60 reference assays sufficed to calibrate the model. Handheld NIRS is a low-cost tool for protein calibration and early-generation screening, while genomic prediction remains suited to parental selection, together supporting a complementary strategy for legume breeding Plain Language SummarySoybean meal is the main protein source for livestock and fish farms in the United States. Because soybean is a summer crop, many Southeastern fields sit idle or grow low-value cover crops in winter. White lupin, a cool-season legume whose seeds are as protein-rich as soybean meal, makes a good complementary winter crop: it yields high-protein grain while serving as a cover crop that fixes nitrogen and frees up soil phosphorus for later crops. In our early-stage lupin breeding program, measuring seed protein by standard lab methods is slow and costly. We built a calibration that lets a handheld scanner estimate protein from light, and compared it with predicting protein from the plants DNA. The scanner gave accurate, low-cost protein screening from only about 40-60 lab tests, while DNA-based prediction remains suited to guiding parent selection. Used together, these tools offer breeders a practical path to develop high-protein white lupin. Core ideasO_LIHandheld NIRS provides screening-grade prediction of white lupin seed crude protein. C_LIO_LISpectra carried more usable protein signal than markers by measuring seed chemistry directly. C_LIO_LINIRS and genomic prediction serve different stages of a white lupin breeding program. C_LIO_LIAbout 40 to 60 reference assays sufficed to calibrate NIRS to near-full accuracy. C_LI

6
Genome-Wide Markers Predict Metribuzin Tolerance in Southern Soft Red Winter Wheat

Sellani, J.; Anzueto, H.; Arcenaux, K.; Price, P. T.; Brown-Guedira, G.; Harrison, S.; DeWitt, N.

2026-07-03 genomics 10.64898/2026.06.28.733875 medRxiv
Top 0.1%
21.6%
Show abstract

Metribuzin is a versatile herbicide effective against various annual grasses and broadleaf weeds found in wheat fields. However, it can cause foliar damage to wheat, impacting plant health and yield. A clearer understanding of the genetic architecture associated with metribuzin tolerance is necessary to guide marker-based breeding strategies. This study evaluated 351 historic Gulf Atlantic Wheat Nursery (GAWN) wheat breeding lines representative of southern US soft red winter wheat (SRWW) germplasm. Field trials were conducted at Winnsboro (WN) and Baton Rouge (BR), Louisiana, in 2016 and 2017. Metribuzin was applied at specific growth stages[DN1.1], and tolerance was assessed based on visual foliar damage. Genomic data from 6,252 filtered single nucleotide polymorphism (SNP) markers were used to estimate narrow-sense heritability, conduct genome-wide association (GWAS), and assess genomic prediction accuracy using genomic best linear unbiased prediction (GBLUP). Broad-sense heritability ranged from 0.54 to 0.69 within environments and reached 0.77 across environments, while narrow-sense heritability ranged from 0.35 to 0.47, indicating moderate additive genetic control. No SNP surpassed the significance threshold, but genomic prediction (GP) showed moderate to strong predictive ability (PA) across environments, with the highest accuracy (r = 0.62) observed between BR17 and WN17. These results indicate that metribuzin tolerance in SRWW is primarily controlled by multiple small-effect loci and that GS provides a more effective breeding strategy than marker-assisted selection for improving tolerance in southern wheat germplasm.

7
Assessing the Role of Marker Density and Minor Allele Frequency on Machine Learning Driven Genomic Selection Accuracy in Grapevine

Francisco, F. R.; de Oliveira, G. L.; Niederauer, G. F.; Fritsche-Neto, R.; Souza, A. P. d.; Furlan, M. F. M.

2026-07-17 genetics 10.64898/2026.07.11.737951 medRxiv
Top 0.1%
19.0%
Show abstract

Although grapevine (Vitis spp.) is among the oldest and most economically significant fruit species globally, its genetic improvement faces major bottlenecks due to long juvenile periods and extended cycles for phenotypic evaluation. In this context, genomic selection (GS) has emerged as an effective alternative to traditional selection, offering a robust framework to optimize breeding programs by significantly reducing generation intervals while enhancing predictive accuracy (PA) in early generations and expected genetic gains (EGGs). Nevertheless, factors such as minor allele frequency (MAF) and population size can significantly affect predictive models, even to the point of making their use unfeasible in breeding programs. In this context, this study evaluated the effect of data dimensionality reduction on GS accuracy by selecting single-nucleotide polymorphisms (SNPs) based on MAF thresholds. The experimental design tested the predictive capacities of four machine learning (ML) algorithms (ElasticNet, K-Neighbors, Support Vector Machine Regression, and XGBoost) alongside the conventional Genomic Best Linear Unbiased Prediction (gBLUP) model. These were validated using three SNP datasets (11,115, 9,494, and 6,100 markers) filtered by MAF levels of 0.05, 0.1, and 0.2 across six genetic traits, and EGGs were compared between conventional breeding and GS via the breeders equation. The results revealed that the ML models exhibited remarkable stability, with no significant differences in PA across the different MAF-based SNP densities, except for berry length, which showed a substantial difference with XGBoost at an MAF of 0.2. Conversely, gBLUP demonstrated high sensitivity to dimensionality reduction, with its performance significantly impacted by MAF filtering across all the traits. These results suggest that compared with traditional GS models that rely on a genomic kinship matrix, ML-based approaches offer greater flexibility in feature reduction. Additionally, compared with chemical traits, morphological traits generally had greater predictive ability. Furthermore, every GS model provided estimated genetic gains superior to traditional breeding, with improvements ranging from an 8.90-fold increase in berry length to a 2.86-fold increase in total soluble solids, confirming that GS integration is promising for enhancing breeding efficiency in grapevines.

8
Multiplex Genome Editing Overcomes Photoperiod Sensitivity in Tropical Maize

Lee, K.; Hampson, E.; Carrillo, R.; Kang, M.; Ghenov, F.; Higa, L.; Du, Z.-Y.; Schoenbaum, G. R.; Yu, J.; Wang, K.; Muszynski, M. G.

2026-08-06 plant biology 10.64898/2026.08.05.742845 medRxiv
Top 0.1%
18.7%
Show abstract

Tropical maize is a rich source of genetic diversity that could enhance temperate maize breeding programs, but its sensitivity to long-day photoperiods, resulting in delayed flowering, limits its widespread use. To overcome this barrier, the Genome Engineering to Sustain Crop Improvement (GETSCI) project used CRISPR/Cas9 to mutate three flowering repressor genes, ZmCCT9, ZmCCT10, and ZmRAP2.7, in the tropical inbred Tzi8. A single sgRNA targeting the first exon of each target gene was combined with an excision cassette carrying the morphogenic genes Babyboom (Bbm) and Wuschel2 (Wus2) to enable efficient transgenic plant regeneration. Transgenic plants carrying frameshift edits in each target gene were recovered, and subsequent crosses produced two non-transgenic genotypes: a double-edited zmcct10, zmrap2.7 line and a triple-edited zmcct9, zmcct10, zmrap2.7 line. Multiple flowering traits were measured for the edited genotypes and unedited Tzi8 inbred in short-day (Hawai i) and long-day (Iowa) field conditions. Both edited genotypes flowered significantly earlier than Tzi8 in both environments. Notably, under long-day conditions, flowering of the two edited lines overlapped with that of the temperate inbred B73, whereas Tzi8 did not. Together, these results demonstrate that targeted, multiplex gene editing can reduce photoperiod sensitivity in a tropical inbred, expanding access to previously untapped genetic diversity for temperate maize improvement.

9
Genetic architecture of soluble arabinoxylan fibre in elite genotypes of bread wheat revealed by genome-wide association analysis

Alabdullah, A. K.; Kosik, O.; Leverington-Waite, M.; Mitchell, R. A.; Prins, A.; Brett, J.; Griffths, S.; Shewry, P. R.; Lovegrove, A.

2026-07-24 genomics 10.64898/2026.07.21.739768 medRxiv
Top 0.1%
18.4%
Show abstract

Dietary fibre intake remains below recommended levels, and increasing fibre content of widely consumed white wheat flour (derived from the starchy endosperm) represents a scalable strategy to improve public health. In wheat starchy endosperm, arabinoxylan (AX) is the dominant fibre component, with the water-extractable (WE) fraction being particularly beneficial to health. We assembled an Elite Fibre Panel (EFP) of 384 elite modern wheat genotypes from UK commercial breeding programmes and quantified the content of WE-AX in wholemeal, as a proxy for soluble AX in white flour, across two UK field environments. Wholemeal WE-AX, showed substantial quantitative variation and moderate genotype-by-environment interaction, with a broad-sense heritability of 0.68. Genome-wide association analyses using 6,791 SNPs identified seven loci associated with WE-AX content. The strongest and most stable effects were mapped to major loci on chromosomes 1B and 6B, previously implicated in AX regulation, while additional loci of smaller and sometimes environment-dependent effect were detected on chromosomes 3A, 5B, and 7A. Favourable alleles increased WE-AX content by [~]5-15% and combined additively. LD-defined intervals contained several high-confidence candidate genes involved in cell-wall biosynthesis, remodelling and post-depositional modification, including PER1, a validated regulator of arabinoxylan cross-linking, together with genes encoding a UTP-glucose-1-phosphate uridylyltransferase, trichome birefringence-like proteins and xyloglucan endotransglucosylase/hydrolases. These findings demonstrate that substantial gains in soluble AX can be achieved by pyramiding favourable alleles already segregating within elite germplasm, providing a practical route for breeding wheat with enhanced dietary fibre content and improved nutritional quality. Key messageMultiple additive loci controlling soluble arabinoxylan content were identified in elite wheat germplasm, enabling marker-assisted breeding for increased dietary fibre in white flour.

10
Comparison of localGEBV and Optimal Haplotype Stacking Fitness Functions using a Novel R Package: HapSelect

Shaffer, W.; Papin, V.; Carter, Z.; Brunner, S. M.; Tong, J.; Villiers, K.; Robinson, H.; Voss-Fels, K.; Hayes, B. J.; Hickey, L.; Dinglasan, E.

2026-07-13 genetics 10.64898/2026.07.08.737160 medRxiv
Top 0.1%
18.0%
Show abstract

Haplotype-based breeding strategies have emerged as promising approaches to maximize long-term genetic gain by identifying complementary parental combinations while maintaining genetic diversity. However, these methods typically require phased genotypes and more intensive workflow pipelines and skillsets. We developed a novel local genomic estimated breeding value (localGEBV) fitness function with similar intent to the optimal haplotype stacking (OHS) framework fitness function and implemented both in the novel R package, HapSelect. Our aim was to evaluate whether phased haplotypes provide additional benefit over the more easily available dosage-based unphased genotypes in highly inbred crops. A subset of bread wheat nested association mapping (NAM) population comprising 444 lines genotyped with 6,054 DArT-Seq markers was analysed. Marker effects were estimated using rrBLUP, localGEBV and haplotype effects were calculated across linkage disequilibrium-defined haploblocks, and genetic algorithms (GA) were used to identify optimal sets of 30 founders using either a localGEBV derived fitness function with unphased, dosage inputs or the OHS fitness function with phased inputs. Selected parental sets were compared with conventional truncation selection (TS) through 150 generations of forward simulation. The OHS fitness function achieved a marginally greater optimized ultimate GEBV than the localGEBV fitness function during GA optimization, with only 18 of the 30 selected founders overlapped between the two methods. Despite these differences, forward simulations demonstrated nearly identical long-term genetic gain for localGEBV and OHS-selected founders, with both approaches outperforming conventional truncation selection by maintaining greater genetic diversity and delaying the genetic plateau. The minimal difference between localGEBV and OHS is likely attributable to the high homozygosity of the population, where localGEBV and haplotype effects are nearly confounded. These results demonstrate that dosage-based localGEBV provides a practical alternative to phased haplotype approaches for parent selection in inbred crops, substantially simplifying genomic workflows while maintaining long-term breeding performance. Future work should evaluate these methods in more diverse inbred populations and outbred species, where great haplotypic diversity may increase the advantage of true haplotype-based optimizations.

11
Leveraging genome-wide association studies and genomic prediction for distinctness, uniformity, and stability (DUS) testing in maize

Daware, A. v.; Hacke, C.; Remay, A.; Starnberger, P.; Schraml, C.; Collonnier, C.; Laurens, F.; Schmid, K. J.

2026-06-12 genetics 10.64898/2026.06.10.731330 medRxiv
Top 0.1%
18.0%
Show abstract

Testing for distinctness, uniformity, and stability (DUS) is a requirement for plant variety registration and based on phenotypic traits, which is time-consuming and sensitive to environmental variation. Advances in genomics allow to complement DUS testing with molecular markers, for which two models in DUS testing were proposed by the Union for the Protection of New Varieties of Plants (UPOV). A use cases was described for maize, but an implementation has been hindered by a lack of suitable markers and validated analytical frameworks. We address these challenges by integrating historical DUS characteristics scores from 352 European hybrid maize varieties with high-density genome-wide single nucleotide polymorphism (SNP) data. Using genome-wide association studies (GWAS), we identified 18 genomic regions and candidate genes associated with 12 DUS characteristics, enabling the development of diagnostic markers consistent with the UPOV model "Characteristic-Specific Molecular Markers". Since most DUS traits are polygenic, we combined GWAS-informed marker selection with XG-Boost-based machine learning to predict notes of DUS characteristics. This approach achieved strong predictive performance across multiple traits (mean accuracy 0.67), demonstrating its potential for managing reference collections under UPOV model "Combining phenotypic and molecular distances in the management of variety collections". Both approaches were validated for two characteristics using independent public USDA-NPGS maize datasets (>1,700 accessions) highlighting the value of public data for method validation. We also identify key limitations of historical DUS data, including imbalanced and sparse trait representation, and discuss mitigation strategies. Despite these constraints, our results demonstrate that molecular markers may improve maize DUS testing, enabling faster, more accurate variety registration and supporting accelerated crop improvement. Key messageHistorical DUS datasets can be used to identify marker-trait associations of DUS characteristics using genome-wide association study (GWAS) and to develop a genomic prediction framework for an accurate prediction of DUS character notes from marker data.

12
Single-kernel near-infrared spectroscopy enables haploid kernel sorting in field and sweet corn using high-oil haploid inducers across diverse donor-inducer combinations

Sharma, S.; Gustin, J. L.; Frei, U. K.; Settles, A. M.; Lübberstedt, T.; Resende, M. F. R.; Hershberger, J.

2026-07-24 plant biology 10.64898/2026.07.23.740370 medRxiv
Top 0.1%
17.1%
Show abstract

Key messageA single-kernel near-infrared reflectance spectroscopy-based sorter can effectively identify haploid kernels for doubled haploid production in field and sweet corn backgrounds. Doubled haploid (DH) technology significantly shortens the breeding cycle for developing homozygous inbred lines in maize (Zea mays). Manual sorting of haploids from a larger bulk of hybrid kernels in an induction cross is a major bottleneck in DH development. Automated systems based on near-infrared (NIR) reflectance spectroscopy can be valuable tools for rapid haploid sorting, provided that sorting accuracy is sufficient for incorporation into the DH process. In this study, we evaluated the accuracy of a custom-built single-kernel NIR (skNIR) sorter for classifying haploid kernels from 12 high-oil haploid induction populations generated from two sweet corn and two field corn donors and four high-oil haploid inducers (HOHIs). We evaluated several general classification models that can be applied without population-specific recalibration or prior genotyping, including models that classified haploids based solely on predicted oil content, as well as multivariate methods that used all wavelengths of the NIR spectra. The highest classification accuracy was obtained using a general multivariate support vector machine (SVM) model. When combined with the two best-performing HOHIs, the general SVM model accurately sorted induction populations from two of the three donor backgrounds crossed with these inducers. Two oil-based methods showed less accurate classification than the multivariate SVM model, due to overlapping oil content distributions across the two kernel classes. Overall, this study demonstrates effective skNIR-based sorting of haploid kernels from diverse induction populations using a single general model. The practical deployment of this instrument in maize breeding programs is discussed.

13
Long-term realized genetic gain and population dynamics under genomic selection in Brazilian cassava germplasm

de Freitas, G. M.; Certuche, D. C. S.; Jannink, J.-L.; De Oliveira, E. J.; Garcia, A. A. F.

2026-07-22 genetics 10.64898/2026.07.18.739356 medRxiv
Top 0.1%
15.3%
Show abstract

Genomic selection has become an important strategy in cassava breeding, enabling faster selection cycles and sustained genetic progress. Despite its widespread adoption, long-term evaluations integrating predictive performance, realized genetic gain, and genetic diversity remain scarce, particularly in clonally propagated crops. We present a comprehensive assessment of genomic selection outcomes in the Brazilian cassava breeding program across four recurrent selection cycles (C0 to C3) implemented between 2011 and 2024, using historical phenotypic and genomic data from 210 multi-environment trials. Predictive ability of genomic best linear unbiased prediction models ranged from low to moderate, depending on the traits genetic architecture and heritability. Prediction accuracies were highest in early cycles (C0 and C1) and showed modest declines in later cycles (C2 and C3). Root yield, shoot yield, plant height, starch content, and dry matter content exhibited stable predictive performance across cycles, with a gradual reduction in RMSE, indicating improved model calibration as training populations expanded. Regression analyses of genomic estimated breeding values revealed significant realized genetic gains for most yield-related traits. In contrast, dry matter content and starch content exhibited small, non-significant negative trends, consistent with known unfavorable genetic correlations with yield. Targeted reductions in plant architecture scores reflected deliberate selection for ideotypes suited to mechanized production systems. At the same time, analyses of genetic diversity revealed a slight decrease in observed heterozygosity, with higher values in the most advanced selection cycle. These results provide an integrated framework for monitoring predictive performance, realized genetic gain, and population genetic dynamics under long-term genomic selection. Collectively, they offer valuable insights into balancing short-term genetic improvement with long-term sustainability and support the development of strategies to optimize selection decisions, breeding planning, and population management in Brazilian cassava breeding programs.

14
Reference-guided comparative genomics of seven Indonesian rice cultivars identifies conserved gene space and trait-associated sequence candidates

Purwestri, Y. A.; Wicaksono, A.; Nurbaiti, S.; Purba, N. T.; Retnaningati, D.; Restiani, R.; Kumalasari, N.; Nuringtyas, T. R.; Handayani, V. D. S.

2026-08-29 genomics 10.64898/2026.08.26.747264 medRxiv
Top 0.1%
14.9%
Show abstract

Indonesian rice cultivars represent valuable genetic resources, yet many remain poorly characterized at the genomic level. Here, we generated 95.40 Gb of PacBio HiFi sequence data from seven Indonesian rice cultivars and constructed cultivar-specific consensus genomes using the telomere-to-telomere Nipponbare reference AGIS1.0. Sequencing coverage ranged from 27.92x to 41.58x, and the resulting consensus genomes spanned 387.93-390.54 Mb, with BUSCO completeness of approximately 98.3-98.5%. OrthoFinder assigned 99.1% of predicted proteins to 40,737 orthogroups, including 27,514 core orthogroups represented across all seven cultivars, indicating a highly conserved predicted gene space within the reference-guided framework. Targeted analysis recovered 278 of 280 cultivar-by-locus combinations representing 40 genes or gene family entries associated with grain pigmentation, nitrogen and amino-acid metabolism, and starch properties. Comparative predicted protein analysis prioritized ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI as candidates for further investigation. Among 269 completed AGIS1.0-anchored promoter comparisons, 159 passed quality-control criteria, whereas 110 were flagged for gene-model, boundary, synteny, or structural concerns. Notably, these flagged comparisons accounted for more than 90% of the alignment-derived sequence variation, emphasizing the importance of rigorous quality control when interpreting apparent promoter divergence. Collectively, these reference-guided genomic resources provide a standardized framework for investigating sequence variation in Indonesian rice germplasm and prioritize testable coding and regulatory candidates for functional validation and future genomics-assisted crop improvement.

15
Whole-genome resequencing identified loci underwent divergent selection and improved local adaptability in groundnut (Arachis hypogaea)

Jahanzaib, M.; He, K.; REHMAN, S.-U.; Ullah, I.; Khurshid, H.; UMER, M. J.; Gangurde, S.; Rasheed, A.; Li, H.

2026-07-18 genomics 10.64898/2026.07.13.738128 medRxiv
Top 0.1%
14.8%
Show abstract

There is an urgent need to expand groundnut genomics knowledge base to improve yield and adaptability in the target environments. Whole-genome sequencing can discover selective sweeps in the genomic regions distinguishing adaptive from non-adaptive germplasm within target environments. When combined with genome-wide association studies (GWAS), this approach can reveal genes underpinning local adaptability and yield advantage. The objective of this study was to establish a genome-wide quantitative framework form identifying genomic regions under selection, with a particular focus on narrowing down regions associated with important yield and adaptive traits. Moreover, validation of elite haplotype distributions in an independent fully sequenced groundnut panel. A panel of 197 groundnut accessions was subjected to whole-genome sequencing and phenotypic evaluation to dissect the collection into adaptive and non-adaptive subsets to uncover the genomic regions under selection. This was then combined with GWAS to uncover genetic variants governing agronomic traits associated with yield and adaptability. The stringent single and multi-trait analysis identified 60 loci for 12 agronomic traits, of which seven loci controlled multiple yield related traits and were pleiotropic. Within our diversity panel, 48 genomic regions showed signs of selection. Among these selective sweeps, six positively selected loci were co-localized with trait associated loci. A large genomic region on chr2 spanning [~]78 Mb was under divergent selection and harbored genes underpinning yield and 20-pod length. A F-box transcription factor, Arahy.37HYKA, on chr9, and an alanine transferase protein gene, Arahy.E9MTVL, on chr12 carried peak SNPs associated with yield and related traits. We further cross-validated our results in another groundnut355 panel, where the corresponding genes within LD blocks showed significant effects on HKW, pod length and pod weight. The genomic resources developed here provide a high-resolution variation map to delineate the genes underpinning yield and developmental traits in groundnut, improved our understanding of the genetic basis of important agronomic traits, and provide a valuable resource for further functional genomics studies and groundnut improvement programs.

16
Early-life stage phenomic prediction of field agronomic traits across breeding cycles in intermediate wheatgrass

Harris, Z. N.; Braley, J.; Cassetta, E.; Crain, J.; DeHaan, L.; Van Tassel, D.; Miller, A.; Rubin, M. J.

2026-08-31 plant biology 10.64898/2026.08.28.747871 medRxiv
Top 0.1%
14.6%
Show abstract

Perennial grains represent a promising frontier for sustainable agriculture, but breeding progress is constrained by the accessibility of genotyping and the difficulty of evaluating complex traits expressed for multiple years after establishment across heterogeneous environments. Phenomic selection may help address these challenges by using inexpensive, scalable, high-dimensional phenotypes collected early in development, although the robustness of such predictions across breeding cycles remains uncertain. Here, we compared genomic selection and phenomic selection across two breeding cycles of Thinopyrum intermedium (intermediate wheatgrass; IWG; Kernza(R)), comprising approximately 2,280 individuals from maternal half-sib families evaluated across multiple field sites and years. We constructed relationship matrices from genomic markers and early-life stage phenomic data, including seed and leaf color (HSV), CropReporter multispectral reflectance and indices, and cycle-specific hyperspectral reflectance sensors. Genomic models provided the strongest predictions on average across all field traits in both cycles. Among phenomic predictors, leaf HSV was consistently the most informative, whereas CropReporter and hyperspectral data showed lower and more trait-dependent performance and seed HSV provided little predictive value. Genomic, leaf HSV, and CropReporter models transferred across breeding cycles with little apparent loss of predictive ability relative to within-cycle validation, demonstrating that their predictive signals were not restricted to a single breeding cycle. Early-life stage leaf HSV emerged as a practical, accessible tool for germplasm thinning and early-stage prioritization in perennial breeding programs. Despite limited similarity among relationship matrices, multi-relationship-matrix models rarely improved prediction beyond the stronger constituent single-relationship-matrix model. Together, these results show that early-life stage phenomic data provide reproducible information about agronomic performance expressed years later, but that predictor complexity and data integration do not guarantee improved prediction.

17
Uncovering genomic regions controlling root quality traits in Cassava (Manihot esculenta Crantz) using different GWAS models

Solarte Certuche, D. C.; Mamedio de Freitas, G.; Jannink, J.-L.; Garcia Morales, C. F.; Sousa Cerqueira, T.; Santos de Santana, B.; Jorge de Oliveira, E.; Garcia, A. A. F.

2026-06-15 genetics 10.64898/2026.06.11.731598 medRxiv
Top 0.1%
12.9%
Show abstract

Cassava is a major staple crop in tropical regions, and improving its root nutritional quality, particularly carotenoid and dry matter content (DMC), remains a central breeding goal. To elucidate the genetic basis of these traits by locating genomic regions associated with them, we analyzed 3,043 cassava clones from the Brazilian Agricultural Research Corporation (Embrapa) breeding program, phenotyped across 188 multi-environment trials conducted from 2011 to 2022 in Brazil. All clones were genotyped using Genotyping-by-Sequencing (27,045 Single Nucleotide Polymorphism - SNPs) and Diversity Arrays Technology - DArTseq (25,923 SNPs). Trait values were estimated using a two-stage mixed model to obtain deregressed BLUPs (Best Linear Unbiased Predictions), and genome-wide association analyses were performed using both the Mixed Linear Model (MLM) and Multi-Locus Mixed Model (MLMM). We detected six significant SNPs consistently associated with carotenoid content and DMC after Bonferroni correction. These SNPs mapped to six candidate genes involved in pathways relevant to root physiology, including Abscisic Acid ABA-related signaling, hydrolase activity affecting carotenoid conversion, fatty-acid biosynthesis within plastids, cell-wall remodeling, and glycolytic energy metabolism. The loci jointly explained 75.56 % of the phenotypic variance for carotenoids and 76.23 % for DMC, with individual SNP effects ranging from [~]17 % to [~]42 % PVE (Proportion of Variance Explained). Broad-sense heritability was H2 = 0.78 for carotenoids and H{superscript 2} = 0.34 for DMC, confirming substantial genetic control and suitability for molecular breeding. Haplotype analyses revealed four superior haplotypes for carotenoids and one key haplotype for DMC, each showing significantly higher trait values compared with other allelic combinations. These haplotypes represent promising targets for marker-assisted selection and genomic selection, with direct applicability for accelerating genetic gain in elite breeding populations. The results provide actionable genomic resources for breeding programs aiming to develop biofortified and high-root quality cultivars and establish a foundation for future multi-omics and functional validation studies.

18
Genome-wide association study of grain iron and zinc concentrations in a diverse CIMMYT wheat panel across contrasting moisture environments

Govindan, V.; Yuan, K.; Dai, Y.; Tarekegn, Z. T.; Lu, L.; Ma, X.

2026-08-06 plant biology 10.64898/2026.08.06.743196 medRxiv
Top 0.1%
12.8%
Show abstract

Micronutrient deficiencies remain a major public-health challenge, and genetic improvement of grain iron and zinc concentrations in wheat offers a sustainable biofortification strategy. We evaluated 563 CIMMYT advanced wheat lines and six checks under restricted irrigation (two irrigations) and well-watered conditions (five irrigations) at Ciudad Obregon, Mexico. Grain iron and zinc concentrations were quantified by energy-dispersive X-ray fluorescence spectrometry. Best linear unbiased estimates were calculated, and genome-wide association analyses were conducted using 8,687 high-quality single-nucleotide polymorphisms and a multi-locus mixed model that accounted for population structure. Five marker-trait associations were detected in at least two datasets. For grain zinc concentration, S3B_811421507 was detected under both irrigation regimes and in the combined analysis, whereas S3B_814372642 was detected under well-watered conditions and in the combined analysis. For grain iron concentration, S2B_73321328 and S4B_20679079 were associated with variation under restricted irrigation and in the combined analysis, while S2B_72162723 was detected under well-watered conditions and in the combined analysis. Individual loci explained 0.34% to 3.36% of phenotypic variance, consistent with the quantitative inheritance of grain micronutrient concentration. The identified alleles provide candidate targets for validation and marker-assisted biofortification breeding, while their environment-dependent effects emphasize the need to evaluate micronutrient traits across contrasting moisture conditions.

19
Haplotypes variations of yellow stripe like (TaYSL) genes are associated with grain iron and zinc contents in wheat (Triticum aestivum L.)

Abbasi, K.; Qayyum, H.; Naseer, S.; Sun, M.; Quraishi, M. A.; Danyal, Y.; Hao, Y.; He, Z.; Rasheed, A.

2026-07-08 plant biology 10.64898/2026.06.17.732851 medRxiv
Top 0.1%
11.8%
Show abstract

The availability of pangenome and resequencing of wheat collections have facilitated the discovery of gene-trait associations in wheat. Yellow stripe-like (YSL) proteins play a key role in the uptake and translocation of metals and yet have not been fully identified and analyzed at the genome-wide level in wheat. In this study, 26 TaYSL genes were identified and divided into four distinct clades, each clade sharing similar domains and motif compositions. Most genes were upregulated under iron deficiency, whereas homoeologs of TaYSL1 were downregulated. Both SNP-based and haplotype-based association studies were used to dissect the role of TaYSLs underpinning grain iron contents (GFeC) and zinc contents (GZnC) in wheat. TaYSL6-2B and TaYSL16-1A haplotypes showed strong association with GFeC, and TaYSL14-6A showed strong association with GZnC in multiple field trials. The distribution of favorable haplotypes in global wheat collection of [~]3000 accessions showed that majority of haplotypes were more prevalent in landraces and winter wheat compared to modern cultivars and spring types, indicating their potential for use in breeding. The combination of favorable haplotypes of three YSL genes associated with GFeC and GZnC were very rare, and most of the wheat accessions has single or double favorable haplotypes. These findings provide the first comprehensive characterization of the TaYSL gene family in wheat and identify significant SNPs and elite haplotypes that can be utilized for genetic improvement and biofortification.

20
A haplotype-based breeding framework for the precise pyramiding of elite QTL alleles: a lettuce case study

Tu, Z.; Luo, G.; Xiao, L.; Wei, M.; Zhang, J.; Wang, X.

2026-08-20 bioinformatics 10.64898/2026.08.12.744550 medRxiv
Top 0.1%
11.7%
Show abstract

The efficient pyramiding of favorable alleles underlying complex traits remains a major challenge in crop breeding as most quantitative trait loci (QTLs) have not been resolved to causal genes, limiting their direct application in marker-assisted breeding. Although haplotypes provide more informative genetic units than individual markers, existing haplotype-based studies have largely focused on genetic interpretation and elite haplotype discovery, whereas computational frameworks for translating haplotypes into breeding decisions remain limited. Here, we developed HAPBDB, a haplotype-guided breeding framework that directly translates regional haplotypes into parental selection, cross design, and elite QTL pyramiding, and applied it to a lettuce genomic breeding panel. HAPBDB accurately reconstructed functional haplotypes at known loci and resolved elite haplotypes for five major QTLs controlling flowering time and yield. Integrating haplotype information across loci enabled systematic identification of accessions carrying complementary elite haplotypes and rational design of crosses that maximized favorable haplotype accumulation while minimizing segregating loci. Experimental validation using QTL-specific molecular markers demonstrated concordance between predicted and observed multi-locus genotypes across all designed F hybrids. Our results demonstrated that regional haplotypes can serve as practical breeding units even when the underlying causal genes remain unknown, thereby enabling the direct utilization of genetically mapped QTLs for precision breeding. By bridging the gap between genomic discovery and practical breeding, HAPBDB provides a practical framework for converting genomic information into breeding decisions and accelerating precision improvement of complex traits.